BetaText: An Event Driven Text Processing and Text Analyzing System
نویسنده
چکیده
BetaText can be described as an event driven pr(xluction system, in which (c~mbinations of) text events lead to certain actions, such as the printing of sentences that exhibit certain, say, syntactic phenomena. %~]e analysis mechanism used allows for arbitrarily complex parsing, but is particularly suitable for finite state i~arsing. A careful investigation of what is actually needed in linguistically relevant text processing resulted in a rather sn%all but carefully chosen set of "elementary actions" to be implemented. 1. Introdnction. The field of c~mputa'tior~][ linguistics seems, roughly speaking, to o~IiprJ.se two rather disjoint subfields, one J.n which the typical researcher predominantly occupies himself witJl problems such as "concordance generation", "backward sorting", "word frequencies" and so on, whereas the prototypic researd]er in tJ~e otJler field has things like "parsing strategies", "semantic representations" on top of his mind. qhis division into almost disjoint subfields is to be regretted, because we all are (or should be) students of one and the same thing language as it is. %~e responsibility for this sad state of affairs can probably be divided equal by the researchers in these two subfields: the "concordance makers" .~cause they seem so entirely ha~)py with rather unsophisticated cx)raputational tools de~eloped a].reac~ in the sixties (and which allow the researcher to look at words or word forms only, and their distribution), and the theoreticians ~yecause they seem so obsessed with the idea of developing their fantastic ir~dels of ]xln(}lage in greater and greater detail, a mode], that at a closer scrutiny is found to c~Dmprise a lexicon of, at best, a couple of hundred words, and cvavering, at best, a couple of hundred sentences or so. No wonder that the researchers in these two canlos thirJ< so little of each other. One way of closing the gap can be to develop niDre sophisticated tools for the investigation of actual texts; there is a need for die theoreticians to test to what extent their models actually cover actual language (and to get impulses from actual language), and there is a need for the "practicioners" to have simple tools for investigating snore complex st[llctures in texts than mere words and word :totals. BetaText is an attempt to provide tools for both those needs. 2. Text events and text oiyerations. BetaText is a system intended both for scientific investigations (or analyses) of texts, and text processing in a i~ore technical sense, such as reformattlng, washing spurious characters away, and so on. Due to the internal organisation of the system, even large texts can [se run at a reasonable cost (of. BroddaKarlsson ±98i). In this section we give some general definitions, and show their consequences for BetaTe xt. i~i elementary (text) event consists of the observation of one specified, concrete string in the text. The systera records sudl an observation through the introduction of a specific internal state (oz through a specific change of the internal state), the internal state being an internal variable that can take arbitrary, positive integral values. /Lrbitrarily chosen states (sets of states, in fact.) can be tied to specific activities (or pro cesses), and each time such a state is intro duced (i.e. the internal state becomes equal to that state) the corresponding process is aeti vated. Such states are called action states. A complex event (or just event, even elementary events can be. cor~lalex in the sense used here) is the c~3mbined result of a sequence of interconnected elementary events, possibly resulting in an action state. In BetaText all this is coi~pletely controlled by a set of prEx~uction rules (cf. Smullyan 196].) of the type~ (, ) -> ( , , , ) where is the string that is to be observed, a condition for applying the rule, viz. that the current inter]lal state belongs to this set; it is via such conditions that the chaining of several elementary events into one con~91ex event is achieved. is a string that is substituted for the observed string (the default is that the original string is retained), is a directive to Che system w'here (in the text) it shall continue the analysis; the default is immediately to
منابع مشابه
Information filtering in high velocity text streams using limited memory: an event-driven approach to text stream analysis
This dissertation is concerned with the processing of high velocity text using event processing means. It comprises a scientific approach for combining the area of information filtering and event processing, in order to analyse fast and voluminous streams of text. In order to be able to process text streams within event driven means, an event reference model was developed that allows for the co...
متن کاملEXTRACTION-BASED TEXT SUMMARIZATION USING FUZZY ANALYSIS
Due to the explosive growth of the world-wide web, automatictext summarization has become an essential tool for web users. In this paperwe present a novel approach for creating text summaries. Using fuzzy logicand word-net, our model extracts the most relevant sentences from an originaldocument. The approach utilizes fuzzy measures and inference on theextracted textual information from the docu...
متن کاملایجاز:یک سامانه عملیاتی برای خلاصهسازی تکسندی متون خبری فارسی
The rapid growth of published documents on the web has created some new requests for processing, classification and information retrieval. So, the use of natural language processing tools has increased around the world. Automatic summarization known as the core of a wide range of text-processing tools such as decision systems, accountability systems, search engines, etc. And always has been inv...
متن کاملDirectional Stroke Width Transform to Separate Text and Graphics in City Maps
One of the complex documents in the real world is city maps. In these kinds of maps, text labels overlap by graphics with having a variety of fonts and styles in different orientations. Usually, text and graphic colour is not predefined due to various map publishers. In most city maps, text and graphic lines form a single connected component. Moreover, the common regions of text and graphic lin...
متن کاملPresenting a method for extracting structured domain-dependent information from Farsi Web pages
Extracting structured information about entities from web texts is an important task in web mining, natural language processing, and information extraction. Information extraction is useful in many applications including search engines, question-answering systems, recommender systems, machine translation, etc. An information extraction system aims to identify the entities from the text and extr...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 1986